Papers with Outcome Reward Models

2 papers
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are prone to hallucination, especially during multihop tasks.
Approach: They propose a hierarchical, erroraware discriminative PRM that classifies math errors at each step and combines finegrained signals to estimate step correctness.
Outcome: The proposed model outperforms the prior best in a new stateof-theart PRMScore of 67.7 on a 400Ksample dataset .
Logical Reasoning with Outcome Reward Models for Test-Time Scaling (2025.emnlp-main)

Copied to clipboard

Challenge: Logical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), but it is under-explored in deductive reasoning.
Approach: They propose to use Chain-of-Thought to generate data using single and multiple samples to train ORMs.
Outcome: The proposed model expands the type of errors covered in the training dataset, covering previously unexplored error types.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations